Papers by Oh Joon Kwon
GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets (2024.emnlp-main)
Copied to clipboard
| Challenge: | Reinforcement learning with human feedback (RLHF) and its offline variant Direct Preference Optimization (DPO) are two of the most important methods for language model (LM) alignment. |
| Approach: | They propose to use a diversity-seeking RL algorithm called GFlowNet-DPO in an offline preference alignment setting to optimize a model's behavior. |
| Outcome: | Empirical results show that the proposed algorithm generates far more diverse responses than the baseline methods and is still relatively aligned with human values in dialog generation and summarization tasks. |
Learning to Embed Multi-Modal Contexts for Situated Conversational Agents (2022.findings-naacl)
Copied to clipboard
Haeju Lee, Oh Joon Kwon, Yunseon Choi, Minho Park, Ran Han, Yoonhyung Kim, Jinhyeon Kim, Youngjune Lee, Haebin Shin, Kangwook Lee, Kee-Eung Kim
| Challenge: | Situated Interactive Multi-Modal Conversations 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs. |
| Approach: | They propose a joint learning approach that integrates visual inputs and performs all four subtasks at once for efficiency. |
| Outcome: | The proposed approach won the 10th Dialog Systems Technology Challenge (DSTC10) . it incorporates visual inputs and performs all four subtasks at once for efficiency . |